[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118809-en":3,"doc-seo-118809-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118809,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Reducing Training Data Needs with Minimal Multilevel Machine Learning (M3L)","Scientific machine learning for chemistry often faces a key bottleneck in acquiring training data rather than in model training itself, even when relying on computation and simulation. To reduce both cost and carbon footprint, this work proposes minimal multilevel machine learning (M3L), which selects training data set sizes by using losses defined at multiple levels of reference data. The objective minimizes prediction error together with computational wall-time acquisition costs. Benchmarks on atomization energies and electron affinities across thousands of organic molecules show substantial reductions in computational effort compared with heuristic multilevel learning.","arXiv :2308 . 11196v1 [physics .chem-ph] 22 Aug 2023  \nReducing Training Data Needs with Minimal Multilevel Machine Learning (M3L)  \nStefan Heinen, 1 Danish Khan, 1, 2 Guido Falk von Rudorff,3, 4 Konstantin Karandashev,5  \nDaniel Jose Arismendi Arrieta,6 Alastair J. A. Price,7, 8 Surajit Nandi,9 Arghya  \nBhowmik,9 Kersti Hermansson,6 and O. Anatole von Lilienfeld 1, 7, 10, 11, 12, 13, ∗  \n1 Vector Institute for Artificial Intelligence, Toronto, ON, M5S 1M1, Canada  \n2 Department of Chemistry, University of Toronto, St. George Campus, Toronto, ON, Canada  \n3 University Kassel, Department of Chemistry, Heinrich-Plett-Str.40, 34132 Kassel, Germany  \n4 Center for Interdisciplinary Nanostructure Science and Technology (CINSaT), Heinrich-Plett-Straße 40, 34132 Kassel  \n5 University of Vienna, Faculty of Physics, Kolingasse 14-16, AT-1090 Wien, Austria  \n6 Department of Chemistry-Ångström Laboratory, Uppsala University, Box 538, SE-75121 Uppsala, Sweden  \n7 Acceleration Consortium, University of Toronto. 80 St George St, Toronto, ON M5S 3H6  \n8 Departments of Chemistry, University of Toronto, St. George Campus, Toronto, ON, Canada  \n9 Department of Energy Conversion and Storage, DTU, Anker Engelunds Vej, DK-2800 Kgs. Lyngby  \n10 Department of Materials Science and Engineering,  \nUniversity of Toronto, St. George campus, Toronto, ON, Canada  \n11 Department of Chemistry, University of Toronto, St. George campus, Toronto, ON, Canada  \n12 Department of Physics, University of Toronto, St. George campus, Toronto, ON, Canada  \n13 Machine Learning Group, Technische Universität Berlin and Berlin  \nInstitute for the Foundations of Learning and Data, Berlin, Germany  \nFor many machine learning applications in science, data acquisition, not training, is the bottleneck even when avoiding experiments and relying on computation and simulation. Correspondingly, and in order to reduce cost and carbon footprint, training data efficiency is key. We introduce minimal multilevel machine learning (M3L) which optimizes training data set sizes using a loss function at multiple levels of reference data in order to minimize a combination of prediction error with overall training data acquisition costs (as measured by computational wall-times) . Numerical evidence has been obtained for calculated atomization energies and electron affinities of thousands of organic molecules at various levels of theory including HF, MP2, DLPNO-CCSD(T), DFHFCABS, PNOMP2F12, and PNOCCSD(T)F12, and treating tens with basis sets TZ, cc-pVTZ, and AVTZF12 . Our M3L benchmarks for reaching chemical accuracy in distinct chemical compound sub-spaces indicate substantial computational cost reductions by factors of ∼ 1.01, 1.1, 3.8, 13.8 and 25.8 when compared to heuristic sub-optimal multilevel machine learning (M2L) for the data sets QM7b, QM9LCCSD(T) , EGP, QM9CCAESD(T) , and QM9CCEASD(T) , respectively. Furthermore, we use M2L to investigate the performance for 76 density functionals when used within multilevel learning and building on the following levels drawn from the hierarchy of Jacobs Ladder: LDA, GGA, mGGA, and hybrid functionals. Within M2L and the molecules considered, mGGAs do not provide any noticeable advantage over GGAs. Among the functionals considered and in combination with LDA, the three on average top performing GGA and Hybrid levels for atomization energies on QM9 using M3L correspond respectively to PW91, KT2, B97D, and τ-HCTH, B3LYP∗ (VWN5), TPSSH.  \nI. INTRODUCTION  \nMachine learning (ML) has revolutionized various scientific fields by leveraging large data sets to extract valuable insights. In 2012 Rupp et al. [1] introduced the first ML approach successfully learning atomization energies. However, the acquisition cost of training data remains a substantial bottleneck, often hindering progress in tackling new challenges. Even more so, data generation in general is still in its infancy: Most of all data has only been collected in the past two decades[2] .  \nAddit","cbCaitpf7vvcudz7","https://ap.wps.com/l/cbCaitpf7vvcudz7","pdf",5829775,1,11,"English","en",105,"# Abstract\n# Introduction\n## Training data acquisition bottleneck in science\n## Computational simulations and HPC cost\n## Multilevel learning (M2L) leading to M3L","[{\"question\":\"What problem does M3L target in scientific machine learning?\",\"answer\":\"M3L targets the bottleneck of training data acquisition cost, which limits progress even when experiments are avoided and computation/simulation are used.\"},{\"question\":\"How does minimal multilevel machine learning (M3L) reduce training data needs?\",\"answer\":\"M3L optimizes training set sizes using losses at multiple levels of reference data to minimize a combination of prediction error and overall acquisition costs measured by computational wall-times.\"},{\"question\":\"What evidence is provided for M3L performance in the document?\",\"answer\":\"The document reports numerical results for atomization energies and electron affinities on thousands of organic molecules, with benchmarks showing cost reductions versus heuristic multilevel learning.\"}]","Reducing Training Data Needs with Minimal Multilevel Machine Learning (M3L) | PDF",1785720367,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"reducing-training-data-needs-with-minimal-multilevel-machine-learning-m3l","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/reducing-training-data-needs-with-minimal-multilevel-machine-learning-m3l/118809/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does M3L target in scientific machine learning?","Question",{"text":76,"@type":77},"M3L targets the bottleneck of training data acquisition cost, which limits progress even when experiments are avoided and computation/simulation are used.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does minimal multilevel machine learning (M3L) reduce training data needs?",{"text":81,"@type":77},"M3L optimizes training set sizes using losses at multiple levels of reference data to minimize a combination of prediction error and overall acquisition costs measured by computational wall-times.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence is provided for M3L performance in the document?",{"text":85,"@type":77},"The document reports numerical results for atomization energies and electron affinities on thousands of organic molecules, with benchmarks showing cost reductions versus heuristic multilevel learning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]