[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122522-en":3,"doc-seo-122522-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122522,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","TRANSFERABILITY OF DATASETS BETWEEN MACHINE-LEARNING INTERACTION POTENTIALS - A PREPRINT","With the rise of Foundational Machine Learning Interatomic Potential (FMLIP) models trained on large datasets, the study addresses how effectively training data can be transferred between different machine-learning forcefield architectures. The work evaluates whether datasets optimized for one ML forcefield algorithm can be reused to fine-tune alternative models, aiming to reduce costly iterative training. Using molecular dynamics on a reactive battery-electrolyte solvent mixture, the approach compares stability and thermodynamic accuracy, tests multiple training configurations and protocols (including MACE), and analyzes transfer behavior for human-designed versus active-learned data, including generalization to unseen molecules.","arXiv :2409 .05590v1 [physics .chem-ph] 9 Sep 2024  \nTRANSFERABILITY OF DATASETS BETWEEN MACHINE-LEARNING INTERACTION POTENTIALS  \nA PREPRINT  \nSamuel P. Niblett 1 , Panagiotis Kourtis2 , Ioan-Bogdan Magdu2 , Clare P. Grey 1 , and Gábor Csányi∗3  \n1Yusuf Hamied Department of Chemistry, University of Cambridge, Lensfield Road, Cambridge, UK  \n2 School of Natural and Environmental Science, Newcastle University, Newcastle upon Tyne, NE1 7RU, UK  \n3Engineering Laboratory, University of Cambridge, Trumpington St and JJ Thomson Ave, Cambridge, UK  \nABSTRACT  \nWith the emergence of Foundational Machine Learning Interatomic Potential (FMLIP) models trained on extensive datasets, the question of how far data can be transferred between different ML architectures has become increasingly important. In this work, we examine the extent to which training data optimised for one machine-learning forcefield algorithm may be re-used to train different models, aiming to accelerate FMLIP fine-tuning and to reduce the need for costly iterative training.  \nAs a test case, we train models of an organic liquid mixture that is commonly used as a solvent in rechargeable battery electrolytes and that plays an important role in degradation processes of these devices, making it an important and representative target for reactive MLIP development. We assess the performance of our models by analysing the stability and thermodynamic accuracy of molecular dynamics trajectories, showing that this is a more stringent test than comparing prediction errors for particular configurations.  \nWe consider several types of training configuration, and several popular machine-learning protocolsnotably the recent MACE architecture, a message-passing neural network designed for high efficiency and smoothness. We demonstrate that simple training sets constructed without any ab initio dynamics simulations are sufficient to produce stable models of molecular liquids that can transfer to multiple liquid compositions. For simple neural-network architectures, further iterative training is required to capture the thermodynamic and kinetic properties of the liquid correctly, but MACE appears to perform well with extremely limited datsets. We find that configurations which are designed by human intuition to correct systematic deficiencies of a model are effectively transferred between algorithms, but that active-learned data that are generated by one MLIP do not typically benefit a different algorithm. As in other tests, MACE shows better performance with transferred activelearned data than traditional neural networks do. Finally, we examine the effect of transferred datasetsize on a model’s ability to generalise to unseen molecules. We find that any training data which improve model performance for the base molecule also improve stability for related unseen molecules, suggesting that trajectory failure modes are connected with chemical structure rather than being entirely system-specific.  \nThese results provide insight into how training set properties affect the behaviour of an MLIP, and practical principles to assist rapid enhancement of training sets for forcefields of molecular liquids.  \nThese approaches may be used in tandem with foundation models to dramatically accelerate the rate at which new chemical systems can be studied by these methods.  \n1 Introduction  \nRecent years have seen an explosion in the field of atomistic simulation for materials and molecular liquids, driven both by the burgeoning importance of electrochemical[1–5] and nanostructured devices[6, 7] to enhance sustainable technology and by the emergence of machine-learning interaction potentials (MLIPs) as a simulation method that  \napproaches quantitative accuracy for comparison with experiment. MLIPs represent an effective compromise between the efficiency of classical molecular dynamics and the accuracy of ab initio quantum calculations,  \nfacilitating simulations of complex molecular systems and material","cbCair9ot1id7QAi","https://ap.wps.com/l/cbCair9ot1id7QAi","pdf",1640817,1,23,"English","en",105,"# Abstract\n# 1 Introduction\n## Background: MLIPs in atomistic simulation\n## Foundational models and fine-tuning needs","[{\"question\":\"What problem does the study focus on regarding ML interaction potentials?\",\"answer\":\"It examines how far training datasets optimized for one machine-learning forcefield algorithm can be reused to train or fine-tune different ML architectures, with the goal of accelerating FMLIP fine-tuning and lowering iterative training cost.\"},{\"question\":\"Why does the study use stability and thermodynamic accuracy as evaluation metrics?\",\"answer\":\"It treats those measures from molecular dynamics trajectories as a more stringent test than comparing prediction errors for individual configurations, because trajectory stability and correct thermodynamic behavior better reflect practical model reliability.\"},{\"question\":\"How do human-designed training configurations compare with active-learned data for transfer between algorithms?\",\"answer\":\"Configurations designed by human intuition to address systematic model deficiencies transfer effectively between algorithms, while active-learned data generated by one MLIP typically does not benefit a different algorithm in the same way. MACE also shows strong performance with transferred active-learned data versus traditional neural networks.\"}]","TRANSFERABILITY OF DATASETS BETWEEN MACHINE-LEARNING INTERACTION POTENTIALS - A PREPRINT | PDF",1785811078,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transferability-of-datasets-between-machine-learning-interaction-potentials-a-preprint","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/transferability-of-datasets-between-machine-learning-interaction-potentials-a-preprint/122522/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study focus on regarding ML interaction potentials?","Question",{"text":75,"@type":76},"It examines how far training datasets optimized for one machine-learning forcefield algorithm can be reused to train or fine-tune different ML architectures, with the goal of accelerating FMLIP fine-tuning and lowering iterative training cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the study use stability and thermodynamic accuracy as evaluation metrics?",{"text":80,"@type":76},"It treats those measures from molecular dynamics trajectories as a more stringent test than comparing prediction errors for individual configurations, because trajectory stability and correct thermodynamic behavior better reflect practical model reliability.",{"name":82,"@type":73,"acceptedAnswer":83},"How do human-designed training configurations compare with active-learned data for transfer between algorithms?",{"text":84,"@type":76},"Configurations designed by human intuition to address systematic model deficiencies transfer effectively between algorithms, while active-learned data generated by one MLIP typically does not benefit a different algorithm in the same way. MACE also shows strong performance with transferred active-learned data versus traditional neural networks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]